Research

The evaluation was anchored to the host review, the CPSP outcome and result-specific RoB 2 assessment.

Before testing the AI, the assessment target had to be made explicit. The host review concerns regional anaesthesia and chronic post-surgical pain (CPSP), and the SWAR uses the RoB 2 tool for individually randomised parallel-group trials.Before testing the AI, the assessment target had to be made explicit. The host review concerns regional anaesthesia and chronic post-surgical pain (CPSP), and the SWAR uses the RoB 2 tool for individually randomised parallel-group trials.

This mattered because RoB 2 is result-specific. A single trial can contribute several outcomes, time points, intervention comparisons and analyses, each of which may carry different risks of bias. Asking an AI to assess ‘the paper’ is therefore methodologically inadequate. The workflow must identify one exact CPSP result, one time point and one pairwise comparison before the first signalling question is answered.


The project also fixed the effect of interest at review level: the effect of assignment to intervention, corresponding to the intention-to-treat estimand. That decision determines which version of Domain 2 is used. It cannot be left for the AI to infer from whichever analysis happens to be easiest to identify in an individual report.


This stage looked like administrative preparation, but it was actually construct control. Without a fixed target result and estimand, two apparently competent assessments may be answering different questions. Agreement statistics would then be tidy, publishable and largely meaningless — the methodological equivalent of measuring two different patients and congratulating the thermometer.